Cross-Tokenizer Challenges in RL
In RL, when the reference model and to-be-trained model have different tokenizers, training with RL algorithms will encounter some problems. Such problems are usually referred to as “cross-tokenizer” (cross-vocabulary) problems.